auxiliary module
Preventing Shortcuts in Adapter Training via Providing the Shortcuts
Adapter-based training has emerged as a key mechanism for extending the capabilities of powerful foundation image generators, enabling personalized and stylized text-to-image synthesis. These adapters are typically trained to capture a specific target attribute, such as subject identity, using single-image reconstruction objectives. However, because the input image inevitably contains a mixture of visual factors, adapters are prone to entangle the target attribute with incidental ones, such as pose, expression, and lighting. This spurious correlation problem limits generalization and obstructs the model's ability to adhere to the input text prompt. In this work, we uncover a simple yet effective solution: provide the very shortcuts we wish to eliminate during adapter training. In Shortcut-Rerouted Adapter Training, confounding factors are routed through auxiliary modules, such as ControlNet or LoRA, eliminating the incentive for the adapter to internalize them. The auxiliary modules are then removed during inference. When applied to tasks like facial and full-body identity injection, our approach improves generation quality, diversity, and prompt adherence. These results point to a general design principle in the era of large models: when seeking disentangled representations, the most effective path may be to establish shortcuts for what should NOT be learned.
TopoReformer: Mitigating Adversarial Attacks Using Topological Purification in OCR Models
Kumar, Bhagyesh, Aravinthakashan, A S, Satyanarayan, Akshat, Gakhar, Ishaan, Verma, Ujjwal
Adversarially perturbed images of text can cause sophisticated OCR systems to produce misleading or incorrect transcriptions from seemingly invisible changes to humans. Some of these perturbations even survive physical capture, posing security risks to high-stakes applications such as document processing, license plate recognition, and automated compliance systems. Existing defenses, such as adversarial training, input preprocessing, or post-recognition correction, are often model-specific, computationally expensive, and affect performance on unperturbed inputs while remaining vulnerable to unseen or adaptive attacks. To address these challenges, T opoReformer is introduced, a model-agnostic reformation pipeline that mitigates adversarial perturbations while preserving the structural integrity of text images. Topology studies properties of shapes and spaces that remain unchanged under continuous deformations, focusing on global structures such as connectivity, holes, and loops rather than exact distance. Leveraging these topological features, T opoReformer employs a topological autoencoder to enforce manifold-level consistency in latent space and improve robustness without explicit gradient regularization. The proposed method is benchmarked on EMNIST, MNIST, against standard adversarial attacks (FGSM, PGD, Carlini-Wagner), adaptive attacks (EOT, BDP A), and an OCR-specific watermark attack (FA W A).
Struct-X: Enhancing Large Language Models Reasoning with Structured Data
Tan, Xiaoyu, Wang, Haoyu, Qiu, Xihe, Cheng, Yuan, Xu, Yinghui, Chu, Wei, Qi, Yuan
Structured data, rich in logical and relational information, has the potential to enhance the reasoning abilities of large language models (LLMs). Still, its integration poses a challenge due to the risk of overwhelming LLMs with excessive tokens and irrelevant context information. To address this, we propose Struct-X, a novel framework that operates through five key phases: ``read-model-fill-reflect-reason'' efficiently enabling LLMs to utilize structured data. It begins by encoding structured data into a topological space using graph embeddings, followed by filling in missing entity information with knowledge retrieval modules, and filtering out irrelevant tokens via a self-supervised module. The final phase involves constructing a topological network with selected tokens to further reduce the total token length for more effective LLM inference. Additionally, Struct-X includes an Auxiliary Module trained to generate prompts, aiding LLMs in analyzing structured data. Extensive experiments on benchmarks, including the knowledge graph question-answer task and the long document reading comprehension task, show that Struct-X notably improves LLM reasoning, demonstrating the effectiveness of structured data augmentation in improving LLM inference with complex input context.
Direct Speech-to-speech Translation without Textual Annotation using Bottleneck Features
Zhang, Junhui, Pan, Junjie, Yin, Xiang, Ma, Zejun
Speech-to-speech translation directly translates a speech utterance to another between different languages, and has great potential in tasks such as simultaneous interpretation. State-of-art models usually contains an auxiliary module for phoneme sequences prediction, and this requires textual annotation of the training dataset. We propose a direct speech-to-speech translation model which can be trained without any textual annotation or content information. Instead of introducing an auxiliary phoneme prediction task in the model, we propose to use bottleneck features as intermediate training objectives for our model to ensure the translation performance of the system. Experiments on Mandarin-Cantonese speech translation demonstrate the feasibility of the proposed approach and the performance can match a cascaded system with respect of translation and synthesis qualities.
Siamese Labels Auxiliary Network(SiLaNet)
Gan, Wenrui, Liu, Zhulin, Chen, C. L. Philip, Zhang, Tong
Auxiliary information attracts more and more attention in the area of machine learning. Attempts so far to include such auxiliary information in state-of-the-art learning process have often been based on simply appending these auxiliary features to the data level or feature level. In this paper, we intend to propose a novel training method with new options and architectures. Siamese labels, which were used in the training phase as auxiliary modules. While in the testing phase, the auxiliary module should be removed. Siamese label module makes it easier to train and improves the performance in testing process. In general, the main contributions can be summarized as, 1) Siamese Labels are firstly proposed as auxiliary information to improve the learning efficiency; 2) We establish a new architecture, Siamese Labels Auxiliary Network (SilaNet), which is to assist the training of the model; 3) Siamese Labels Auxiliary Network is applied to compress the model parameters by 50% and ensure the high accuracy at the same time. For the purpose of comparison, we tested the network on CIFAR-10 and CIFAR100 using some common models. The proposed SilaNet performs excellent efficiency both on the accuracy and robustness.
Better Resource Usage through Biomimetic Symbiotic Principles for Host and Derivative Product Synthesis
Davis, Matthew Louis Turner (Texas A&M University) | McAdams, Daniel Arthur (Texas A&M University) | Wadia, Anosh Porus (Texas A&M University)
In recent years, numerous methods to aid designers in conceptualizing new products have been developed. These methods intend to give structure to a process that was, at one time, considered to be a purely creative exercise. Resulting from the study, implementation, and refinement of design methodologies is the notion that both the structure of the development process and the structure of the developed product are key factors in creating value in a firm’s product line. With respect to the latter key factor, product architecture, but more specifically, modular product architecture has been the subject of much study. This research is focused on two tasks: advancing the notion of a modular product architecture in which modules can be incorporated into a product ‘post-market,’ and creating a method that aids designers leverage knowledge of natural symbiotic relationships to synthesize these post-market modules. It adds to prior work by first, defining the terms ‘derivative product’ and ‘host product’ to describe the post-market module and the product that the module augments, respectively. Second, by establishing three guidelines that are used to assess the validity of potential derivative products, giving the newly termed host and derivative product space defined boundaries. And lastly, by developing a 7-step, biomimetic-based methodology that can be used to create derivative product concepts (post-market modules). By using this methodology, the engineered products are designed on symbiotic principles found in nature.